Skip to content

optimize qwen3.8-flash-next - #185

Merged
hjh0119 merged 16 commits into
modelscope:mainfrom
hjh0119:qwen4-0903
Sep 7, 2026
Merged

optimize qwen3.8-flash-next#185
hjh0119 merged 16 commits into
modelscope:mainfrom
hjh0119:qwen4-0903

Conversation

@hjh0119

@hjh0119 hjh0119 commented Sep 3, 2026

Copy link
Copy Markdown
Collaborator

No description provided.

hjh0119 added 12 commits August 26, 2026 17:23
main landed the same feature independently in modelscope#174, so qwen4_exp.py /
ple.py / qsa_indexer.py / hyper_connection_gated.py came out as add/add
conflicts. Resolved to this branch's versions -- they are supersets:
  select_mask              -> selection_as_mask (+ selection_as_token_indices,
                              select_token_indices_thd for the sparse kernel)
  _set_ple_ngram_embedding -> fill_table_from_hf / export_table_to_hf
                              (offload-aware, and reduces the offload flag
                              across pp before gating pp collectives)
  _warn_qsa_fallback_once  -> dropped with the fallback path itself

Kept main's ple_seed instead of this branch's hardcoded _PLE_SEED: the
parser derives it from text_config.seed (defaulting to 1234), which stays
configurable and still avoids the vLLM config-pollution issue. Wired it
through Qwen4ExpTextNGramEmbedding with the same fail-loud treatment as
eos_token_id / split_ngram_parts.

Also picks up modelscope#175 (packed sequence length handling for mcore 0.16.0/0.16.1).
hjh and others added 2 commits September 7, 2026 16:28
# Conflicts:
#	src/mcore_bridge/model/gpts/qwen4_exp.py
#	src/mcore_bridge/tuners/lora.py
@hjh0119
hjh0119 marked this pull request as ready for review September 7, 2026 09:31
@hjh0119
hjh0119 merged commit d874528 into modelscope:main Sep 7, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants